Go top
Paper information

The MERIT Dataset: Modelling and efficiently rendering interpretable transcripts

I. de Rodrigo, A. Sánchez-Cuadrado, J. Boal, A.J. López López

Pattern Recognition Vol. 172, nº. Part B, pp. 112502

Summary:

This paper introduces the MERIT Dataset, a multimodal, fully labeled dataset of school grade reports. Comprising over 400 labels and 33k samples, the MERIT Dataset is a resource for training models in demanding Visually-rich Document Understanding tasks. It contains multimodal features that link patterns in the textual, visual, and layout domains. The MERIT Dataset also includes biases in a controlled way, making it a valuable tool to benchmark biases induced in Language Models. The paper outlines the dataset’s generation pipeline and highlights its main features and patterns in its different domains. We benchmark the dataset for token classification, showing that it poses a significant challenge even for SOTA models.


Spanish layman's summary:

MERIT es un dataset sintético de documentos (imagen+texto+maquetación) para Document AI. Generado por un pipeline abierto, permite entrenar y evaluar VLMs en extracción de información, interpretabilidad por embeddings y análisis de sesgos.


English layman's summary:

MERIT is a synthetic document dataset (image+text+layout) for Document AI. Built with an open pipeline, it enables training and evaluation of VLMs on key information extraction, embedding-based interpretability, and bias analysis.


Keywords: Synthetic Dataset; Multimodal Dataset; Visually-rich Document Understanding; Vision-Language Models


JCR-JIF Impact Factor and WoS quartile: 9,100 - Q1 (2025)

DOI reference: DOI icon https://doi.org/10.1016/j.patcog.2025.112502

Published on paper: April 2026.

Published on-line: September 2025.



Citation:
I. de Rodrigo, A. Sánchez-Cuadrado, J. Boal, A.J. López López, "The MERIT Dataset: Modelling and efficiently rendering interpretable transcripts", Pattern Recognition, Vol. 172, nº. Part B, pp. 112502, April 2026. [Online: September 2025] doi: 10.1016/j.patcog.2025.112502

    Research topics:
  • Deep Learning for Industrial Process and Asset Optimization
    Research groups:
  • Instituto de Investigación Tecnológica (IIT)
    ODS:
  • Goal 9: Industry, innovation and infrastructure